AI Learning Series · Perception — the engineer's lens

AI as an Observability Tech Stack

Once the model becomes a runtime dependency of your software, you need to observe it the way we observe distributed systems — because it fails in ways your dashboards were never designed to see.

Foundations
→
Systems
→
Practice
→
Observability (Perception)

01 The Big Picture

You can't manage what you can't see. Distributed systems forced the industry to invent Distributed Tracing, RED metrics, and SLOs. AI is now the same kind of dependency — a network endpoint your product's correctness rests on — with the classic observability pathologies amplified.

Your agent, copilot, or RAG pipeline is a component calling a flaky, semi-opaque service over the network. Treat it exactly like that, and three amplifications follow:

Emergent behavior

The "service" contains a model whose behavior emerges from billions of parameters you did not write and cannot enumerate. System diagrams end at your harness boundary (doc 23); past that line, behavior can only be measured, never read.

Non-determinism

Same input, different output — every time, by design. A passing test case proves one sample, not the system. Percentage-based truth is forced on you, which changes the math (§03).

Plausible wrongness

Failures come back as fluent, well-formed, confident responses. HTTP 200, p99 fine, no exception thrown — and the answer is quietly wrong. The worst failure mode of all, because it is invisible to every classic signal.

Why traditional APM isn't enough: APM assumes success is binary. For a payment service, latency and error codes cover nearly everything. For an AI feature, success has degrees — a response can be statistically right, format-valid, and on time, and still be contextually wrong: citing the retrieved document that doesn't quite answer the question, calling the right tool with subtly wrong arguments, escalating when it should have deflected. The critical signals for an AI system live above the infrastructure layer, in semantics and quality — so we need observability planes stacked on top of tracing, not instead of it.

🔑
Positioning: 'doc 13 built evaluations — those are your unit tests and regression suites, run deliberately before/at deploy. AI observability is production monitoring: the same measurement math applied continuously, on live traffic, with alerting. Different discipline, same math.

02 The What — Four Observability Planes

AI observability is not one tool; it is a stack of four planes, each answering a different question, each reading from the trace the previous plane emitted.

plane 1 · telemetry — "what happened, at what price?"Per-request spans: prefill/decode split (doc 06), TTFT & ITL (doc 24), token counts per class plane 1 · telemetry (cont.)Cost telemetry per request: fresh-input / cached-input / output token classes priced 1× / ~0.1× / 3–5× (doc 07) · structured logs of every tool call and agent turn (doc 23) · context-window payload snapshots — hash or redact: this plane carries PII plane 2 · quality — "was what happened any good?"Online evals as continuous monitors: LLM-as-judge sampling at k% of traffic · exact-match / rubric checks where ground truth exists · sampled human review as calibration anchors (doc 25's calibrated decision readout, JEV) · price-weighted canary prompts replayed hourly plane 3 · semantics — "is the population still the same?"Embedding drift on the input distribution: cosine-similarity shift and cluster drift with doc 11 math · retrieval hit-rate as a service indicator for RAG · hallucination-likelihood signals as proxies — never a single number of truth plane 4 · governance — "did it stay inside the rules?"Audit logs · redaction verification · permission-denied counts from the doc 26 permission hooks · cost-quota enforcement per task type and per tenant

Reading the four planes as one stack: the telemetry plane records facts (tokens, spans, dollars — cheap to collect, but silent about quality). The quality plane grades samples of those recorded facts, at judge prices you must budget (§03). The semantics plane watches the distribution rather than individual cases — cheap embedding math catches a traffic shift long before quality metrics move. The governance plane turns several of those signals into obligations: the audit log is built from the same spans; the redaction check runs over the same payload snapshots; the quota enforcer reads the same cost counters. Everything joins on one trace ID. That single design decision — one trace, four projections — is what makes this an observability stack rather than four disconnected dashboards.

03 The Why — SLI/SLO Math for AI

The Golden Signals translate — with an AI accent on two of them:

Golden SignalClassic versionAI version
LatencyRequest duration percentilesTTFT and ITL percentiles separately — a request can be fast to start and slow to stream (doc 24)
TrafficRequests / secTokens / sec per class + requests / sec — cost and load live in the token axis
Errors5xx rate, exceptionsDeflection/escalation rate · assert-based refusals · hallucination-likelihood — plus classic errors; plausible wrongness hides under a 200
SaturationCPU, queue depthRate-limit headroom · KV-cache pressure (doc 06, serving docs) · batch queue depth · cache-hit ratio (your doc 07 lever)

The sampled error budget. Classic error budgets tolerate exact counts because every request is instrumented and success is observable per request. For agent tasks, "task succeeded" is often only decidable by grading — and grading costs model calls. So you grade a sample, and infer the true rate from n sampled task outcomes: p̂ = k ⁄ n failed tasks out of n total. Sampling means the number carries uncertainty, and that uncertainty must be sized — this is the same confidence math as doc 13, now online.

Wilson interval (90% CI), n = 200 sampled agent tasks, k = 14 failures → p̂ = 0.07, z = 1.645:

center = (p̂ + z²/2n) ⁄ (1 + z²/n) = (0.07 + 2.706/400) ⁄ 1.01353 ≈ 0.0768 half-w = z·√(p̂(1−p̂)/n + z²/4n²) ⁄ (1 + z²/n) ≈ 0.0300 90% CI ≈ [0.047, 0.107] → width ≈ 6 points

Two consequences: (1) with n = 200 you can wrap a 7% error rate usefully tightly — a ±3-point band — but you cannot detect a 1-point regression with this sample size; that needs n ≈ 1,000+ samples per point. (2) The error budget algebra is unchanged: budget burn = (n_failed ⁄ n_total) − (1 − SLO). If your agent-task SLO is 95% success and the sample says 93.0%, you burned 2 points of budget this period — but only know that within ±3 points, so SLA decisions on barely-breached budgets must wait for the next window or a bigger sample.

Judge cost, made explicit. The quality plane's dominant cost is the judge. Judging traffic at rate k with a judge that reads about the same tokens as the judged task costs roughly:

judge spend ≈ k × (judge_cost_per_task ⁄ task_cost_per_task) × total_AI_spend

Numeric example: 10,000 tasks/day at ₹1.00/task = ₹10,000/day of product spend. Sampling k = 5% and judging with a small, cheap model at ₹0.25/task: 0.05 × 0.25 × 10,000 = ₹125/day — about 1.25% overhead, essentially free observability. Judge with a frontier model at ₹4/task and k = 20%, and you are paying ₹8,000/day: an 80% tax. The sampling rate and the judge price class together are one instance of the price-class thinking of doc 07 applied to monitoring.

Cost as a first-class signal. Don't just record spend — alert on it. A burn-rate alert pages when hourly token spend (per task type, per model tier) deviates beyond a few σ from a trailing baseline: a runaway agent loop that retries a failing tool twenty times shows up as spend-sigma long before any user complains. Cost telemetry is not the finance team's report; it is your saturation + runaway-failure channel in one signal.

04 How It Works — One Request, Traced End-to-End

Step through a single agent request and watch it become telemetry, then quality signals, then a dashboard update — the planes stack on one trace.

1 · user message trace root begins here agent span (router) plan → decide tool · tokens: 1,850 fresh prefill 12K tok @1× · TTFT 380 ms child span · tool: search_kb(0.9 s) 4 docs in · hit-rate logged · top-k=4 child span · tool: run_query(1.4 s) in+out tokens: 610 + 2,300 · decode ITL 14 ms annotation · cache HIT — prefix 12K tok billed @ ~0.1× (doc 07) · cost telemetry shows it output stream · 520 output tokens ITL p50 = 12 ms · p99 = 38 ms · $ per token class logged online-eval sampler ← this trace (5%) judge reads span snapshot graded 0.6 → low-score bucket dashboard / SLO panel budget burn ↑ · drift alert armed

Why all seven boxes belong in one picture: each plane's data was emitted by the request itself. The cache-hit annotation is one attribute on the prefill span; the ITL percentiles are derived from the stream timestamps; the sampler scores the stored 5% of spans; the panel aggregates score samples against the SLO. One trace ID threads all four planes — no stitching, no guessing.

05 Reference Architecture — The Instrumented Pipeline

One trace per agent turn, OpenTelemetry-style spans per stage — naming follows the emerging OpenTelemetry GenAI semantic conventions conceptually; check the current spec before implementing, as of this writing the conventions are still stabilizing.

ONE TRACE · trace_id = user request http.server / app span session_id · user_tier · feature flag gen_ai.agent span agent name · task type · turn # policy check permission hook · redact gen_ai.tool spans tool name · args-hash · result retrieval span query · chunk ids · hit-rate context snapshot hash/redacted window payload gen_ai.completion prefill/decode · TTFT · ITL COLLECTOR / PIPELINE → three simultaneous sinks metrics backend (percentiles, burn rates) · log store with trace_id join (evidence) · eval queue (sampler picks 5%, judges, writes scores back by trace_id) DASHBOARD · user-visible SLO panel quality SLIs beside latency/cost · error-budget gauge drift panel: embedding shift, cluster movement ALERTER SLI breach · budget burn · spend-σ · drift gate routes: on-call → #ai-incidents · quality → product owners

Per-stage attributes, and the alert each stage can credibly fire:

Span nameKey attributesAlert that could fire
gen_ai.completionmodel, temperature, prompt_tokens {fresh, cached}, completion_tokens, ttft, itl_p95TTFT p99 breach · token-spend σ · cache-hit ratio drop (doc 07 lever)
gen_ai.tooltool name, args-hash, outcome, duration, retries, costtool p99 latency · retry/timeout spike · permission-denied surge (doc 26 hooks)
retrievalquery, chunk_ids, hit_rate, overlap_score (doc 11 math)retrieval hit-rate SLI below floor → RAG quality incident
eval.judgetrace_id of graded sample, judge model, score, rubric versionscore distribution shift vs. calibrated band; judge–human disagreement ↑
drift.monitorcosine-sim shift, cluster drift %, n computed over (doc 11)input-distribution drift gate — change-coupled releases flagged before faults surface

The drift alert deserves emphasis: it is the cheapest early-warning signal you own. Quality metrics move only as fast as samples accumulate; embedding statistics over the input distribution move in near-real time, and you can compute them from already-logged artifacts at negligible cost — every alert above can name-check OpenTelemetry-style conventions without depending on their exactness.

06 The Evals-as-Monitors Mapping

Doc 13's eval stages are the same instruments listed in a different schedule. This table is the Rosetta stone between the two disciplines.

Eval / monitor stageClassic observability analogueWhen it runsWhat it protects
Unit checkTest in CIEvery change you makeThe grader / rubric itself, on a single known case
Regression suiteTest suite + CI gatePre-merge / pre-deployQuality on a fixed set of mined past failures
Nightly suiteSynthetic checkNightly batch over datasetA slow-moving aggregate baseline — the "yesterday is fine" reference point
Canary prompts (hourly replay)Canary deployment / synthetic probeHourlyProvider-side change without a version bump; early TTFT / loss signs; quality trend
Online judge (sampled k%)Black-box probe / continuous SLO trackingAlways-on, sampledThe error budget itself (§03) — where p̂, CIs, and burn rates meet live traffic
Sampled human reviewCalibrating a test environment vs. realityWeekly sample, graded by peopleThe judge's own calibration — the JEV axis of doc 25 — against a ground-truth anchor
User feedback (👍/👎)Satisfaction survey / CSATContinuous, biasedSemantically the coarsest signal — direction, never magnitude; treat as a trend channel

Read the table once and the thesis becomes operational: eval discipline and monitoring discipline are one measurement practice on two schedules — same rubrics, same graders, same dataset lineage, all inherited from doc 13.

07 Engineering Takeaways

Dashboards to build first. Cost per interaction by task type (with cached/fresh/output split — the doc 07 price table as columns) · tool-call p99 latency and retry rates · escalation rate (agent hands off to a human) · cache-hit ratio. These four cover ~80% of early incidents with almost no infra, because they're all derived from spans you already want for debugging.
Alert routing. Distinguish infrastructure-class pages (provider throttling, rate-limit saturation, spend-σ) from quality-class tickets (score-distribution shift, hit-rate sag, user-thumb trend). Route on-call to infra, product owners to quality; page only on budget burn.
Incident review for agents: "what did the agent see?" The retrieval trace and the document set it retrieved are the evidence — store it. Most AI incidents aren't "the model is dumb": they're "the model was fed a stale doc / a redaction that broke a chunk / an empty search result and then confabulated." The payload snapshot from the telemetry plane is the entire case file.
Runbook: hallucination incident. (1) Triage: representative failed traces — grep trace_id → payload snapshots, retrieval spans, judge scores. (2) Check the drift gate — did the input distribution shift? Compare cluster centroids for the affected task type against the pre-incident baseline (doc 11 math). (3) Check retrieval hit-rate — did chunking / index content change (doc 11)? (4) Check judge calibration — is this a real quality drop or a miscalibrated sensor reading? (5) Blast-radius: fraction of traffic under the affected task type. (6) Fix: revert prompt/docs/change, re-run canaries, then post-incident: an added eval case sampled from the failed span — that failure is a permanent regression test.
The cache-hit ratio is your heat-gauge on prompt hygiene (doc 07). If it drops from ~0.92 to ~0.7, prompt structure regressed — timestamps in the prefix, reordered tools. Cost telemetry catches it as pure bug-class signal; very few teams even monitor it.

08 Mental Models

The model is a new kind of dependency — observe it like a flaky third-party service

You don't have its source, you can't debug its internals, and the provider's SLA guarantees valid responses, not specific answers. Every hard-won discipline for third-party dependencies — health checks, synthetic probes, contract checks, exit plans, evidence retention — applies. The model is simply the dependency where the CI test suite is weakest, so production observation has to carry more weight.

Evals = tests in CI · online monitors = production alerting

The exact recast of doc 13's four-piece eval under an operational load factor. Tests run before launch over a fixed dataset you own; online monitors run after launch, on traffic you cannot curate, at prices you must budget. Same measurement math — sampling, CIs, calibration — different schedule and different assumptions. Keeping them in one codebase (same rubrics, same graders, same judge prompts) is why the mapping table of §06 actually works when it matters.

09 Common Misconceptions

"We have logs — we're observable." Logs ≠ traces ≠ evals. Logs give you what your harness wrote, unjoined; traces give you causality and latency structure through the turn loop (doc 23); evals give you the grade your logs can never self-supply. Without a judge reading the stored traffic at some sample rate, "observability" goes blind above exactly the layer where your AI product's risks live.

"LLM-as-judge is a metric, so its score is the truth." A judge is a noisy sensor, not an oracle. Its score is another model's estimate, with its own drift, truncation blindness, length bias, and self-preference. This is exactly why sampled human review (doc 25's calibrated-readout connection) is on the table: you calibrate the sensor against people, you do not replace the people with the sensor.

"Cost monitoring is finance's job." Cost is a first-class SLO input. Token spend is simultaneously your saturation signal (what's spending, at what rate, against whose quotas) and your runaway-failure channel (a stuck agent loop shows up as a spend-σ anomaly before any user traces it). A monitoring posture that only routes cost to finance discards the operational signal inside it.

"We'll add observability once the tooling is mature." It is cheap to start — one trace ID at the model-call boundary, a few structured logs, one sampled judge — and AI systems drift silently with every model update and document-set change. The "before" baseline cannot be reconstructed after the fact: instrument at launch, or your first quality incident doubles as the start of your data collection.